Random forests for feature selection in QSPR Models - an application for predicting standard enthalpy of formation of hydrocarbons
نویسندگان
چکیده
BACKGROUND One of the main topics in the development of quantitative structure-property relationship (QSPR) predictive models is the identification of the subset of variables that represent the structure of a molecule and which are predictors for a given property. There are several automated feature selection methods, ranging from backward, forward or stepwise procedures, to further elaborated methodologies such as evolutionary programming. The problem lies in selecting the minimum subset of descriptors that can predict a certain property with a good performance, computationally efficient and in a more robust way, since the presence of irrelevant or redundant features can cause poor generalization capacity. In this paper an alternative selection method, based on Random Forests to determine the variable importance is proposed in the context of QSPR regression problems, with an application to a manually curated dataset for predicting standard enthalpy of formation. The subsequent predictive models are trained with support vector machines introducing the variables sequentially from a ranked list based on the variable importance. RESULTS The model generalizes well even with a high dimensional dataset and in the presence of highly correlated variables. The feature selection step was shown to yield lower prediction errors with RMSE values 23% lower than without feature selection, albeit using only 6% of the total number of variables (89 from the original 1485). The proposed approach further compared favourably with other feature selection methods and dimension reduction of the feature space. The predictive model was selected using a 10-fold cross validation procedure and, after selection, it was validated with an independent set to assess its performance when applied to new data and the results were similar to the ones obtained for the training set, supporting the robustness of the proposed approach. CONCLUSIONS The proposed methodology seemingly improves the prediction performance of standard enthalpy of formation of hydrocarbons using a limited set of molecular descriptors, providing faster and more cost-effective calculation of descriptors by reducing their numbers, and providing a better understanding of the underlying relationship between the molecular structure represented by descriptors and the property of interest.
منابع مشابه
QSPR models to predict thermodynamic properties of some mono and polycyclic aromatic hydrocarbons (PAHs) using GA-MLR
Quantitative Structure-Property Relationship (QSPR) models for modeling and predicting thermodynamic properties such as the enthalpy of vaporization at standard condition (ΔH˚vap kJ mol-1) and normal temperature of boiling points (T˚bp K) of 57 mono and Polycyclic Aromatic Hydrocarbons (PAHs) have been investigated. The PAHs were randomly separated into 2 groups: training and test sets. A set o...
متن کاملImproving QSPR models for predicting standard enthalpy of formation with a hybrid approach for feature selection
متن کامل
Capturing the Crystal: Prediction of Enthalpy of Sublimation, Crystal Lattice Energy, and Melting Points of Organic Compounds
Accurate computational prediction of melting points and aqueous solubilities of organic compounds would be very useful but is notoriously difficult. Predicting the lattice energies of compounds is key to understanding and predicting their melting behavior and ultimately their solubility behavior. We report robust, predictive, quantitative structure-property relationship (QSPR) models for enthal...
متن کاملToward Quantitative Structure-Property Relationships for Charge Transfer Rates of Polycyclic Aromatic Hydrocarbons.
Quantitative structure-property relationships (QSPRs) have been developed and assessed for predicting the reorganization energy of polycyclic aromatic hydrocarbons (PAHs). Preliminary QSPR models, based on a combination of molecular signature and electronic eigenvalue difference descriptors, have been trained using more than 200 PAHs. Monte Carlo cross-validation systematically improves the per...
متن کاملPrediction of boiling point and water solubility of crude oil hydrocarbons using sub-structural molecular fragments method
The quantitative structure–property relationship (QSPR) method is used to develop the correlation between structures of crude oil hydrocarbons (80 compounds) and their boiling point and water solubility. Sub-structural molecular fragments (SMF) calculated from structure alone were used to represent molecular structures. A subset of the calculated fragments selected using stepwise regression (fo...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
عنوان ژورنال:
دوره 5 شماره
صفحات -
تاریخ انتشار 2013